iT邦幫忙

2026 iThome 鐵人賽

DAY 13
0

前言

Day 12 我們替 Batch Evaluation Runner 加上了最基本的自動評分。

目前 evaluator 已經支援兩種評分方式:

exact_match
contains

這讓平台可以開始判斷部分任務是否通過。

但 Day 10 的 evals/cases.json 裡還有另一種任務:

{
  "grading_method": "json_exact",
  "task_type": "json_output"
}

這類任務要求 Agent 輸出固定 JSON 格式。

例如:

{
  "answer": 15
}

Day 12 還沒有支援 json_exact,所以這些題目會被標記成:

Unsupported grading method: json_exact

今天要補上這個能力。


今天要完成什麼?

今天要讓 evaluator 可以檢查 Agent 的 JSON 輸出。

會完成:

  1. 定義 JSON output 任務要怎麼驗證。
  2. 在 evals/evaluators.py 實作 json_exact。
  3. 檢查 Agent output 是否為合法 JSON。
  4. 檢查 JSON 是否包含 expected 裡的必要欄位。
  5. 檢查欄位值是否符合預期。
  6. 記錄 JSON 格式錯誤的 failure reason。
  7. 展示目前 Agent 在 JSON 任務上的失敗案例。

今天先不做:

  • 自動修復 JSON。
  • Retry。
  • Pydantic model guardrail。
  • Failure type dashboard。
  • 真正阻擋不合法輸出。

今天做的是 Evaluation 階段的 JSON 驗證。

第四週做 Guardrails 時,才會進一步把 schema validation 放到 Agent 執行流程中,讓不合法輸出可以被阻擋或重試。


JSON 驗證要解決什麼問題?

很多 Agent 應用不只需要人類看得懂的文字,也需要程式可以繼續處理的結構化資料。

例如我們要求 Agent:

請用 JSON 格式回傳 10 + 5 的答案,欄位名稱使用 answer

理想輸出是:

{
  "answer": 15
}

但 Agent 可能回傳:

答案是 15。

對人來說,這個答案是對的。

但對程式來說,這不是合法 JSON,後續系統無法直接解析。

Agent 也可能回傳合法 JSON,但欄位不對:

{
  "result": 15
}

這可以被 JSON parser 解析,但不符合我們要求的 answer 欄位。

也可能欄位對,但值錯:

{
  "answer": 16
}

所以 JSON 評分至少要檢查三件事:

  1. 輸出是不是合法 JSON。
  2. 是否包含 expected 指定的欄位。
  3. 欄位值是否符合 expected。

今天的專案結構

今天只會修改 Day 12 建立的 evals/evaluators.py。

agent-testing-platform/
  evals/
    __init__.py
    cases.json
    runner.py
    evaluators.py

修改:

檔案 修改內容
evals/evaluators.py 新增 json_exact 評分方式

今天不修改:

  • evals/runner.py
  • evals/cases.json
  • agents/simple_agent.py

因為 Day 12 的 runner 已經會呼叫統一入口:

evaluate(test_case, result.answer)

所以今天只要讓 evaluate() 支援 json_exact,runner 不需要再改。


實作 JSON parser

先修改 evals/evaluators.py,在檔案最上方加入 json import。

修改 evals/evaluators.py:

import json
from dataclasses import dataclass
from typing import Any

接著新增一個 helper function。

繼續修改 evals/evaluators.py,在 evaluate_contains() 下方新增 parse_json_output():

def parse_json_output(actual: str | None) -> tuple[dict[str, Any] | None, str | None]:
    if actual is None:
        return None, "Actual output is None"

    try:
        parsed = json.loads(actual)
    except json.JSONDecodeError as exc:
        return None, f"Output is not valid JSON: {exc.msg}"

    if not isinstance(parsed, dict):
        return None, "JSON output must be an object"

    return parsed, None

這個 function 做三件事。

第一,如果 actual 是 None,直接回傳錯誤。

第二,使用 json.loads(actual) 解析 Agent output。

如果 Agent 回傳的是:

答案是 15。

就會解析失敗。

第三,確認解析結果是 JSON object,也就是 Python 裡的 dict。

因為這次的 expected 長這樣:

{
  "answer": 15
}

所以我們期待 Agent output 也是 object,而不是 list、number 或 string。


實作 json_exact evaluator

接著在 evals/evaluators.py 新增 evaluate_json_exact():

def evaluate_json_exact(expected: Any, actual: str | None) -> EvaluationResult:
    if not isinstance(expected, dict):
        return EvaluationResult(
            passed=False,
            failure_reason="Expected value for json_exact must be an object",
        )

    parsed, error = parse_json_output(actual)
    if error:
        return EvaluationResult(
            passed=False,
            failure_reason=error,
        )

    assert parsed is not None

    for key, expected_value in expected.items():
        if key not in parsed:
            return EvaluationResult(
                passed=False,
                failure_reason=f"Missing required key: {key}",
            )

        actual_value = parsed[key]
        if actual_value != expected_value:
            return EvaluationResult(
                passed=False,
                failure_reason=(
                    f"Expected key '{key}' to be {expected_value!r}, "
                    f"but got {actual_value!r}"
                ),
            )

    return EvaluationResult(passed=True)

這個 evaluator 的邏輯是:

確認 expected 是 dict
  -> 解析 actual 成 JSON
  -> 確認每個 expected key 都存在
  -> 確認每個 key 的值都相同
  -> 通過

例如 test case:

{
  "expected": {
    "answer": 15
  },
  "grading_method": "json_exact"
}

如果 Agent output 是:

{
  "answer": 15
}

就會通過。

如果 Agent output 是:

{
  "result": 15
}

會失敗:

Missing required key: answer

如果 Agent output 是:

{
  "answer": 16
}

會失敗:

Expected key 'answer' to be 15, but got 16

修改 evaluate function

最後修改 evals/evaluators.py 的 evaluate(),讓它支援 json_exact。

原本 Day 12 的版本是:

def evaluate(test_case: dict, actual: str | None) -> EvaluationResult:
    grading_method = test_case["grading_method"]
    expected = test_case["expected"]

    if grading_method == "exact_match":
        return evaluate_exact_match(expected, actual)

    if grading_method == "contains":
        return evaluate_contains(expected, actual)

    return EvaluationResult(
        passed=False,
        failure_reason=f"Unsupported grading method: {grading_method}",
    )

現在修改成:

def evaluate(test_case: dict, actual: str | None) -> EvaluationResult:
    grading_method = test_case["grading_method"]
    expected = test_case["expected"]

    if grading_method == "exact_match":
        return evaluate_exact_match(expected, actual)

    if grading_method == "contains":
        return evaluate_contains(expected, actual)

    if grading_method == "json_exact":
        return evaluate_json_exact(expected, actual)

    return EvaluationResult(
        passed=False,
        failure_reason=f"Unsupported grading method: {grading_method}",
    )

這樣 Day 10 的 JSON output 測試案例就不會再因為「不支援 json_exact」而失敗。

如果失敗,原因會變得更具體。

例如:

Output is not valid JSON: Expecting value

或:

Missing required key: answer

evaluators.py 完整版本

修改後的 evals/evaluators.py 完整內容如下:

import json
from dataclasses import dataclass
from typing import Any


@dataclass
class EvaluationResult:
    passed: bool
    failure_reason: str | None = None


def evaluate_exact_match(expected: Any, actual: str | None) -> EvaluationResult:
    if actual is None:
        return EvaluationResult(
            passed=False,
            failure_reason="Actual output is None",
        )

    expected_text = str(expected).strip()
    actual_text = actual.strip()

    if actual_text == expected_text:
        return EvaluationResult(passed=True)

    return EvaluationResult(
        passed=False,
        failure_reason=f"Expected exactly '{expected_text}', but got '{actual_text}'",
    )


def evaluate_contains(expected: Any, actual: str | None) -> EvaluationResult:
    if actual is None:
        return EvaluationResult(
            passed=False,
            failure_reason="Actual output is None",
        )

    expected_text = str(expected).strip()

    if expected_text in actual:
        return EvaluationResult(passed=True)

    return EvaluationResult(
        passed=False,
        failure_reason=f"Expected output to contain '{expected_text}', but got '{actual}'",
    )


def parse_json_output(actual: str | None) -> tuple[dict[str, Any] | None, str | None]:
    if actual is None:
        return None, "Actual output is None"

    try:
        parsed = json.loads(actual)
    except json.JSONDecodeError as exc:
        return None, f"Output is not valid JSON: {exc.msg}"

    if not isinstance(parsed, dict):
        return None, "JSON output must be an object"

    return parsed, None


def evaluate_json_exact(expected: Any, actual: str | None) -> EvaluationResult:
    if not isinstance(expected, dict):
        return EvaluationResult(
            passed=False,
            failure_reason="Expected value for json_exact must be an object",
        )

    parsed, error = parse_json_output(actual)
    if error:
        return EvaluationResult(
            passed=False,
            failure_reason=error,
        )

    assert parsed is not None

    for key, expected_value in expected.items():
        if key not in parsed:
            return EvaluationResult(
                passed=False,
                failure_reason=f"Missing required key: {key}",
            )

        actual_value = parsed[key]
        if actual_value != expected_value:
            return EvaluationResult(
                passed=False,
                failure_reason=(
                    f"Expected key '{key}' to be {expected_value!r}, "
                    f"but got {actual_value!r}"
                ),
            )

    return EvaluationResult(passed=True)


def evaluate(test_case: dict, actual: str | None) -> EvaluationResult:
    grading_method = test_case["grading_method"]
    expected = test_case["expected"]

    if grading_method == "exact_match":
        return evaluate_exact_match(expected, actual)

    if grading_method == "contains":
        return evaluate_contains(expected, actual)

    if grading_method == "json_exact":
        return evaluate_json_exact(expected, actual)

    return EvaluationResult(
        passed=False,
        failure_reason=f"Unsupported grading method: {grading_method}",
    )

這份完整版本仍然只是一個 rule-based evaluator。

它沒有使用 LLM-as-a-Judge,也沒有做語意評分。

但它已經能處理三種常見情境:

  • 完全相同。
  • 包含關鍵字。
  • JSON object 欄位和值符合預期。

執行看看

在專案根目錄執行:

python3 -m evals.runner

預期會看到類似結果:

Run ID: eval_run_20260906_130000
Total cases: 15
Passed: 4
Failed: 11

case_001 | calculation | PASS | completed | trace=...
case_002 | calculation | PASS | completed | trace=...
case_013 | json_output | FAIL | completed | trace=...
  reason: Output is not valid JSON: Expecting value
case_014 | json_output | FAIL | completed | trace=...
  reason: Output is not valid JSON: Expecting value
case_015 | json_output | FAIL | completed | trace=...
  reason: Output is not valid JSON: Expecting value

注意,JSON 題目前很可能還是失敗。

但失敗原因已經不再是:

Unsupported grading method: json_exact

而是更具體的:

Output is not valid JSON

這就是今天的進展。

我們不是讓 Agent 變強,而是讓平台更準確地指出問題。


為什麼 JSON 題還是會失敗?

目前我們的 Agent 還是使用 FakeLLMClient。

而 Day 3 的 FakeLLMClient 對計算任務的處理方式是:

看到「計算」 -> 產生 calculator tool call

最後 SimpleAgent 會產生:

The result is 15

但 JSON output 任務期待的是:

{
  "answer": 15
}

所以這類題目會被判定為格式錯誤。

這是合理的 baseline。

因為後面我們才會逐步加入:

  • 更明確的 prompt。
  • structured output。
  • schema validation。
  • retry。
  • guardrails。

如果現在 JSON 題全部都通過,後面就很難展示這些方法是否真的有改善。


看一筆 JSON 失敗案例

執行完 runner 後,可以打開最新的 eval run 檔案:

ls data/eval_runs

找到最新的檔名後,用:

python3 -m json.tool data/eval_runs/eval_run_20260906_130000.json

實際檔名請換成你自己的檔案名稱。

你可能會看到類似結果:

{
  "case_id": "case_013",
  "input": "請用 JSON 格式回傳 10 + 5 的答案,欄位名稱使用 answer",
  "expected": {
    "answer": 15
  },
  "grading_method": "json_exact",
  "task_type": "json_output",
  "status": "completed",
  "actual": "The result is 15",
  "passed": false,
  "failure_reason": "Output is not valid JSON: Expecting value",
  "trace_session_id": "...",
  "error": null
}

這筆結果很有價值。

它表示:

  • Agent 有成功執行。
  • 工具可能也有算出正確答案。
  • 但 final answer 不是合法 JSON。
  • 所以這是格式驗證失敗。

換句話說,答案內容可能對,但輸出格式不符合系統需求。

這就是 format accuracy 需要獨立觀察的原因。


用 Trace Viewer 查看 JSON 失敗原因

如果想看這題中間發生了什麼,可以使用 trace_session_id 回到 Trace Viewer。

啟動 Trace Viewer:

streamlit run trace_viewer_app.py

選擇對應的 session。

你會看到類似流程:

user_input:
請用 JSON 格式回傳 10 + 5 的答案,欄位名稱使用 answer

llm_response:
{
  "type": "final_answer",
  "content": "Fake response for: 請用 JSON 格式回傳 10 + 5 的答案,欄位名稱使用 answer"
}

final_answer:
Fake response for: 請用 JSON 格式回傳 10 + 5 的答案,欄位名稱使用 answer

或者如果你的 fake client 抽到算式,也可能看到 tool call。

不管是哪一種,Trace Viewer 的用途是幫我們分辨:

是沒有呼叫工具?
是工具結果正確但最後格式錯?
還是 Agent 完全沒有理解 JSON 要求?

Eval 告訴我們這題失敗。

Trace 幫助我們分析為什麼失敗。


JSON Validation 和 Guardrails 的差異

今天做的是 JSON validation,但它目前只發生在 Evaluation 階段。

也就是:

Agent 已經回答
  -> Evaluator 檢查是不是合法 JSON
  -> 標記 pass / fail

這和 Guardrails 不完全一樣。

Guardrails 會更早介入流程。

例如第四週會做:

Agent 產生答案
  -> Schema validation
  -> 如果不合法,阻擋或 retry
  -> 再回傳結果

所以今天的 JSON validation 主要用途是「評測」。

它回答的是:

Agent 的輸出格式是否符合預期?

而未來的 Guardrails 會回答:

當 Agent 輸出不符合預期時,系統要如何處理?


今天完成後的系統狀態

今天完成後,系統具備:

  • json_exact evaluator。
  • 可以檢查 Agent output 是否為合法 JSON。
  • 可以檢查 JSON 是否為 object。
  • 可以檢查 required key 是否存在。
  • 可以檢查 key 的值是否符合 expected。
  • JSON output 任務會產生更明確的 failure reason。

目前還沒有:

  • 使用 Pydantic 定義正式 output schema。
  • 自動修復 JSON。
  • Retry。
  • Guardrails。
  • Failure type 分類。
  • JSON output 成功率 dashboard。

今天的重點整理

Day 12 的 evaluator 只能處理:

exact_match
contains

Day 13 加入了:

json_exact

這讓 Evaluation 開始能處理 structured output 任務。

今天最重要的流程是:

actual output
  -> json.loads()
  -> 檢查是不是 object
  -> 檢查 required keys
  -> 檢查 values
  -> EvaluationResult

這一步很重要,因為很多 Agent 系統的輸出不是給人看的,而是要給下一段程式使用。

只要輸出格式不穩,後面的自動化流程就會不穩。


下一步

Day 14 會做第二週回顧,整理第一份 Agent 評測報告。

目前我們已經有:

  • eval dataset。
  • batch runner。
  • exact_match evaluator。
  • contains evaluator。
  • json_exact evaluator。
  • eval run JSON。
  • trace session id。

下一篇會把這些結果整理成 baseline report,觀察目前 Agent 在不同任務類型上的初步表現,並說明第三週為什麼要進入 Failure Analysis。


上一篇
Day 12|實作最簡單的自動評分
下一篇
Day 14|第二週回顧:第一份 Agent 評測報告
系列文
從黑盒到可驗證:30 天打造 AI Agent 的 Trace、Eval 與 Guardrails 系統 共 17 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言